optimized hardware streaming architecture for the cnn (Xilinx Inc)
Structured Review

Optimized Hardware Streaming Architecture For The Cnn, supplied by Xilinx Inc, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/optimized+hardware+streaming+architecture+for+the+cnn/optimized+hardware+streaming+architecture+for+the+cnn/pmc07288095-410-10-16
Average 90 stars, based on 1 article reviews
Images
1) Product Images from "Real-Time Energy Efficient Hand Pose Estimation: A Case Study"
Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study
Journal: Sensors (Basel, Switzerland)
doi: 10.3390/s20102828
Figure Legend Snippet: The architecture of Convolutional Neural Network (CNN)-based hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.
Techniques Used: Activation Assay
Figure Legend Snippet: Design process overview; the first box illustrates the software-level design phase, while the other boxes illustrate the hardware related design phase. The first stage is quantization-aware training (QAT) in which we decrease the CNN memory demand as well as the computation time on the hardware. The second stage is hardware streaming architecture (HSA) where the underlying hardware structure is designed for the CNN. In system integration (SI) stage, the programmable logic PL and the processing system PS are brought together and the interface with the memory is configured through the hardware system integration sub-stage. Furthermore, the on-Chip Software is developed for preprocessing and interfacing with the peripherals.
Techniques Used: Software
Figure Legend Snippet: Streaming Architecture. Each CNN layer is mapped into a hardware block, and the hardware blocks are connected to each others via stream channels. The bitwidth of each stream is shown on this figure.
Techniques Used: Blocking Assay
Figure Legend Snippet: Hardware system integration; AXI-Lite interface provides the interconnection between the PS and the PL. DMA module is integrated in the PL. This module is responsible for converting the memory mapped input to AXI stream CNN input, as well as converting the AXI stream CNN output to a memory mapped output.
Techniques Used:
Related Articles
Activation Assay:Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study Article Snippet: and effectively quantizing the network model parameters, which resulted in a significantly compressed model for a negligible decrease in accuracy. .. Afterwards, we provided an Software:Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study Article Snippet: and effectively quantizing the network model parameters, which resulted in a significantly compressed model for a negligible decrease in accuracy. .. Afterwards, we provided an Blocking Assay:Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study Article Snippet: and effectively quantizing the network model parameters, which resulted in a significantly compressed model for a negligible decrease in accuracy. .. Afterwards, we provided an |